Papers with multimodal Large Language Models
AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness (2025.acl-long)
Copied to clipboard
| Challenge: | Existing models that assess mLLMs on harmful meme understanding are inaccurate and lack accuracy. |
| Approach: | They propose a framework that adaptively probes the reasoning capabilities of mLLMs . their framework systematically reveals the varying performance of different target mllms a . |
| Outcome: | The proposed framework systematically reveals the performance of different target mLLMs. |
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation metrics for image captioning are primarily designed for short captions and are not suitable for long captions. |
| Approach: | They propose an automatic evaluation metric for long captions developed within a novel LLM-Hybrid-as-a-Judge framework. |
| Outcome: | The proposed metric outperforms existing metrics and achieves superhuman performance on LongCap-Arena. |
Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data (2024.findings-acl)
Copied to clipboard
Yanda Li, Chi Zhang, Gang Yu, Wanqi Yang, Zhibin Wang, Bin Fu, Guosheng Lin, Chunhua Shen, Ling Chen, Yunchao Wei
| Challenge: | OpenAI's GPT-4 has demonstrated remarkable multimodal capabilities, but specific mechanics of GPT4 remain unknown. |
| Approach: | They propose a data collection methodology that synchronously synthesizes images and dialogues for visual instruction tuning. |
| Outcome: | The proposed method improves on ten commonly assessed models and provides greater flexibility compared to existing methods. |
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation approaches focus on mLLMs’ detection accuracy for binary classification tasks, which often fail to reflect the in-depth interpretive nuance of harmfulness across diverse contexts. |
| Approach: | They propose an agent-based arena-style evaluation framework that provides context-aware and unbiased assessment for mLLMs’ understanding of multimodal harmfulness. |
| Outcome: | The proposed framework reduces evaluation biases of judge agents and provides unbiased comparisons of mLLMs’ abilities to interpret multimodal harmfulness. |
Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in multimodal Large Language Models (LLMs) have significantly enhanced the automation of medical image analysis, but still suffer from hallucinations and clinically significant errors. |
| Approach: | They propose a grounding fixation strategy that integrates radiologist eye fixations and bounding box annotations into the LLM prompting framework. |
| Outcome: | The proposed model improves performance without retraining across domain-specific and general-purpose models and achieves an 87.3% clinical average performance. |